🌻 What VAR tells us about high-stakes qualitative judgement

7 Aug 2026

I am a moderate football fan, I follow Man United, Bosnia and Herzegovina and England a bit. Like several million other people I was watching on 11 July when Djed Spence went down in the Norway penalty area in extra time of the World Cup quarter final, and Clément Turpin pointed at the spot, and I did jump up and shout a bit.

Then Turpin jogged over to the monitor and gave it back. Spence had put his leg across in front of the defender, maybe so Spence made the contact, so no penalty. Watching it again I still could not have told you and you have to also somehow guess intention and what who was thinking about what and the causal properties of their causal beliefs. England went through anyway, thanks to Bellingham.

What I learned is that a World Cup quarter final hung on one person's judgement on who caused a collision, made in about (an endless number of) seconds in front of eighty thousand people and a global audience, all of whom had could see the same replays. And noone suggested that this was an unsound way to decide it.

The rules are written down (in this case)#

Take hand-ball, where the law has been chewed over for years by people who are certainly not stupid. This is Law 12 of the Laws of the Game, 2026/27 edition, which are IFAB's rather than FIFA's, me neither. It is an offence if a player:

touches the ball with their hand/arm when it has made their body unnaturally bigger. A player is considered to have made their body unnaturally bigger when the position of their hand/arm is not a consequence of, or justifiable by, the player's body movement for that specific situation.

It's well written. Note, not a consequence of, or justifiable by, the player's body movement, for that specific situation. Justifiable to whom? You are being asked to imagine what a defender's arm was for, half a second ago at a full sprint, and then say whether that was a reasonable place for an arm to be based on your intuition of the player's intentions.

The one part of it they could make measurable, they did: for handball purposes the upper boundary of the arm is in line with the bottom of the armpit. Precise and checkable, but mostly not what the arguments are about.

There is a second moment of judgement too. The VAR protocol does not let the video officials referee the match again. They may only intervene where the on-field decision was a "clear and obvious error" or a serious missed incident, and the final decision is always the referee's. So some benighted person in a windowless room has to judge how wrong a colleague's judgement was, on a scale with no units, in about forty seconds, live on television.

No algorithm can do any of this at the moment. So several experienced people look at the pictures, argue a bit, and one of them decides. And the world accepts it more or less.

At which point I start thinking about qualitative research methods, like you do.

We do this everywhere when something really matters#

A dozen assorted strangers sit through six weeks of testimony and decides who did what to whom, with what intention, and whether the doing caused the harm, and then a person goes to prison for eleven years. A panel of judges in The Hague decides whether an order given in one capital caused a massacre in another. A select committee takes evidence for a year and reports on whether the policy caused the collapse. A board sits down and works out whether the accident rate fell because of the new procedure, or the new manager, or the mild winter, or providence, or nothing much at all.

These are all human perceptions of causation, formed out of testimony and documents by people arguing with each other. We give those judgements the most authority we have got. We take away somebody's liberty on the strength of one. Without much use of what you might call scientific data.

And then you get to social research#

Somewhere between the courtroom and the seminar room the same faculty gets demoted, without anyone announcing it. What people say caused what turns into data about the speaker rather than evidence about the world. If a claim cannot be turned into a measurement it goes in the interesting-but-soft drawer, in the smaller wing of the department round the back, next to the person who does poetry. Lovely people. No budget.

Call it the Backyard Assumption: a causal judgement that cannot be measured is a charming second best, to be indulged politely until the real methods turn up. In some workshops, no-one will challenge you for talking like that, though that depends what evaluation universe you inhabit.

The case for the prosecution#

Let me put it at its strongest, because it is a good case and there is no sport in demolishing a poor straw man.

First, people are bad at judging causes. Two experts read the same case and disagree. The same expert reads it again on a Thursday and disagrees with both of them. Perfectly confident every time. See (Kahneman 2011).

Secondly, people are even worse at judging their own causes. Nisbett and Wilson showed in 1977 that we confabulate our own reasons, fluently and sincerely. Ask a farmer why their yields went up and you get a story. Quite a good story. And the respondent knows who is paying for the interview.

Thirdly, a judgement nobody can check is not evidence. "I read the interviews and formed a view" leaves the reader two options, belief or disbelief, and neither of those is an argument. A number can at least be recomputed and attacked in public.

Fourthly, and this is the one that ought to bother us most, none of it floats free of power. Days after Bosnia went out to the USA, with Folarin Balogun sent off for a foul on Tarik Muharemović, the President of the United States rang the President of FIFA to say the red card had been "horrible". FIFA duly discovered Article 27 of its Disciplinary Code, the ban was suspended, and Balogun played Belgium in the round of 16. UEFA said it crossed a line. Bosnia had gone home.

And what counts as a foul in one league is a shrug in another, everyone knowing that interpretation is much of a muchness within a country and something else across the border. The evaluator judging the causal claims is paid by the funder, trained in one tradition, usually working in English, and decides whose account is evidence and whose is anecdote. A method built on human judgement is a method in the hands of whoever is holding the whistle.

Stack the four of them up and you get an unreliable faculty, applied to unreliable testimony, written up in a form nobody can audit, by somebody who answers to somebody. If that were all I had to go on, I would want numbers too.

Taking them in turn#

On the first. That research compares judgement against a formula. Where a formula is to be had, use it! Nobody here is arguing for vibes. But there is no formula for whether an arm was justifiable by the player's body movement for that specific situation, and the match restarts in ninety seconds. Wittgenstein's point: no rule can tell you how to apply itself. So the only question is whether the judgement is made well or badly. Turned the other way up, all that research is a design brief: write the rule down, show everyone the same evidence, name who decides, keep the record. Which is the VAR room.

On the second. Nisbett and Wilson were studying introspection, i.e. people reporting on their own mental processes. "They rebuilt the road, so I could get my tomatoes to market before they spoiled" is nothing of the kind. It is a claim about events out in the world, of the sort a witness gives in court, and it can be set against other witnesses and the delivery records. Courts met unreliable testimony with cross-examination rather than exclusion. Not perfect. Not nothing.

On the third, which is the good one. It is right. But checkability is a property of the procedure, and has little to do with numbers. The VAR room gives you a published rule, evidence the decider and the audience see at the same moment, a named decider, a stated threshold for overruling them, a record, and a fortnight of argument by grown adults quoting the same clause at each other. Not a single number in the building. What the measurement rule really demands is show me your rule, show me your evidence, and let me check whether you applied the one to the other. Measurement is one way of meeting that. It was never the only one.

Evaluation has had its own Law 12 for years, and we call it a rubric: write down what good would look like, in words, before anybody goes near the evidence, then show your reasoning against it in public (King et al. 2013). There are versions built specifically for causal contribution claims (Aston 2019). Nobody who has sat on a moderation panel would call that soft.

On the fourth, which I cannot really answer. Power gets into judgement, and anybody who tells you their procedure has fixed that is selling you something, probably a toolkit. But notice how we all know about the Trump call: a named decision-maker, a rule that had to be cited by number, a published outcome, an opponent who could appeal. It is on the record and stays there.

Would a number have saved us? Who picks the indicator, who defines the category, whose ministry supplies the figure. Power does its best work early, back where the categories get fixed, and by the time a number reaches the page the argument is finished and invisible. Judgement leaves fingerprints. Measurement wipes them off and calls it objectivity. None of which makes any of it fair. It makes it contestable, which is the only sort of fairness anyone has ever actually been offered.

And while we are here, the unresolved calls are not the embarrassment they look like either. When two pundits go at each other for a fortnight over a handball, they are arguing about the same incident, quoting the same clause, and each of them can tell you what would change their mind. Gallie called this an essentially contested concept, i.e. one whose application people go on arguing about without either side being the least bit muddled about what the argument is over. Ask two labour statisticians whether a man who did four hours of paid work last week counts as employed and you get the identical row, except that it ends in a number which hides the fact that the row ever took place.

What happened when they did hand it to the machine#

Eight days before the Spence penalty, in the last 32, Croatia went out to Portugal. Gonçalo Ramos had put Portugal 2-1 up deep into stoppage time, and then Joško Gvardiol equalised, and for a few seconds Croatia were level and Luka Modrić's last World Cup was still alive. Espen Eskas was sent to the monitor. The goal was disallowed, because the sensor inside the Adidas Trionda ball had registered a faint touch off Igor Matanović on the way in, possibly off his hair, which put Mario Pašalić offside before the assist. FIFA say the data are sound, and they are probably right, some sense or other. Or are they?

The touch was not visible on any camera as far as I could see. The telly showed was a graphic of the sensor trace, a sort of heartbeat, which is presumably not lying. Croatia fans threw bottles on the pitch, which is a form of peer review I would not recommend but can understand.

This is the one part of football we have automated, and look at what it cost to do it. Offside is a partly geometry problem, so of course a machine help with the geometry. But to settle this one the officials had to accept a kind of evidence that does not agree with what you can see with your own eyes. Arguably the meaning of "touched" should be what humans can normally sense, not one from physics. I think they are wrong to do that.

So the lesson is not that machines are bad and humans are good. It is that measuring the measurable does not diminish your problem of making judgements about the unmeasurable. Duh.

Which is more or less what we are up to in Causal Map#

A coder reads a passage and decides one thing: is this person claiming that A influenced B? The rule is written down and it is short, because we code only bare causation and leave out all the fascinating extra gubbins the coder is itching to add. The decision is stored against the quotation that prompted it, so anybody can rewind the tape. A second coder goes through the same passages, and where the two of them disagree they are disagreeing about one sentence, not about a whole report.

I have coded links I would not want read back to me in court. That is rather the point: you can read them back to me. So the numbers on a causal map are counts of judgements and we should keep saying so in public, because they have roughly the standing of a tournament's log of VAR decisions. Then compare the alternative, which is a set of causal claims spread over four paragraphs of very good prose, based on evidence the reader cannot see, made by a person who is not named, against a rule that nobody ever wrote down.

An AI does the first pass in a lot of our work now. The rule is written, the quotation stays attached to the link, and somebody still has to say "I vouch for this".

We trust this faculty with prison sentences and treaties and World Cups, and then ask it to apologise for itself in an evaluation report. Show your working, by all means, that is what the whole apparatus is for. But sending human causal perception out to the backyard is a strange thing to do with the only instrument we have ever really had.

Or have I got that wrong?

References

Aston (2019). Contribution Rubrics.

Kahneman (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.

King, McKegg, Oakden, & Wehipeihana (2013). Evaluative Rubrics: A Method for Surfacing Values and Improving the Credibility of Evaluation. Journal of MultiDisciplinary Evaluation, 9, 11--20.